NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Odin: Learning to Optimize Operation Unit Configuration for Energy-efficient DNN Inferencing

https://doi.org/10.23919/DATE64628.2025.10993275

Narang, Gaurav; Doppa, Janardhan Rao; Pande, Partha Pratim (March 2025, IEEE)

Free, publicly-accessible full text available March 31, 2026
Heterogeneous Manycore In-Memory Computing Architectures

https://doi.org/10.1145/3676536.3697138

Ogbogu, Chukwufumnanya; Narang, Gaurav; Joardar, Biresh Kumar; Doppa, Janardhan Rao; Pande, Partha Pratim (October 2024, ACM)

Full Text Available
HuNT: Exploiting Heterogeneous PIM Devices to Design a 3-D Manycore Architecture for DNN Training

https://doi.org/10.1109/TCAD.2024.3444708

Ogbogu, Chukwufumnanya; Narang, Gaurav; Joardar, Biresh Kumar; Doppa, Janardhan Rao; Chakrabarty, Krishnendu; Pande, Partha Pratim (November 2024, IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems)

Full Text Available
Dataflow-Aware PIM-Enabled Manycore Architecture for Deep Learning Workloads

https://doi.org/10.23919/DATE58400.2024.10546730

Sharma, Harsh; Narang, Gaurav; Doppa, Janardhan Rao; Ogras, Umit; Pande, Partha Pratim (March 2024, IEEE)

Full Text Available
TEFLON: Thermally Efficient Dataflow-aware 3D NoC for Accelerating CNN Inferencing on Manycore PIM Architectures

https://doi.org/10.1145/3665279

Narang, Gaurav; Ogbogu, Chukwufumnanya; Doppa, Janardhan_Rao; Pande, Partha_Pratim (August 2024, ACM Transactions on Embedded Computing Systems)

Resistive random-access memory (ReRAM)-based processing-in-memory (PIM) architectures are used extensively to accelerate inferencing/training with convolutional neural networks (CNNs). Three-dimensional (3D) integration is an enabling technology to integrate many PIM cores on a single chip. In this work, we propose the design of athermallyefficient dataflow-aware monolithic 3D (M3D)NoC architecture referred to asTEFLONto accelerate CNN inferencing without creating any thermal bottlenecks.TEFLONreduces the Energy-Delay-Product (EDP) by 42%, 46%, and 45% on an average compared to a conventional 3D mesh NoC for systems with 36-, 64-, and 100-PIM cores, respectively.TEFLONreduces the peak chip temperature by 25Kand improves the inference accuracy by up to 11% compared to sole performance-optimized SFC-based counterpart for inferencing with diverse deep CNN models using CIFAR-10/100 datasets on a 3D system with 100-PIM cores.
more » « less
Uncertainty-Aware Online Learning for Dynamic Power Management in Large Manycore Systems

https://doi.org/10.1109/ISLPED58423.2023.10244486

Narang, Gaurav; Ayoub, Raid; Kishinevsky, Michael; Doppa, Janardhan Rao; Pande, Partha Pratim (August 2023, IEEE)

Full Text Available
DYNAMIC POWER MANAGEMENT IN LARGE MANYCORE SYSTEMS: A LEARNING-TO-SEARCH FRAMEWORK

https://doi.org/10.1145/3603501

Narang, Gaurav; Deshwal, Aryan; Ayoub, Raid; Kishinevsky, Michael; Doppa, Janardhan Rao; Pande, Partha Pratim (July 2023, ACM Transactions on Design Automation of Electronic Systems)

The complexity of manycore System-on-chips (SoCs) is growing faster than our ability to manage them to reduce the overall energy consumption. Further, as SoC design moves towards 3D-architectures, the core's power density increases leading to unacceptable high peak chip temperatures. In this paper, we consider the optimization problem of dynamic power management (DPM) in manycore SoCs for an allowable performance penalty (say 5%) and admissible peak chip temperature. We employ a machine learning (ML) based DPM policy, which selects the voltage/frequency (V/F) levels for different cluster of cores as a function of the application workload features such as core computation and inter-core traffic etc. We propose a novel learning-to-search (L2S) framework to automatically identify an optimized sequence of DPM decisions from a large combinatorial space for joint energy-thermal optimization for one or more given applications. The optimized DPM decisions are given to a supervised learning algorithm to train a DPM policy, which mimics the corresponding decision-making behavior. Our experiments on two different manycore architectures designed using wireless interconnect and monolithic 3D demonstrate that principles behind the L2S framework are applicable for more than one configuration. Moreover, L2S-based DPM policies achieve up to 30 energy-delay product savings and reduce the peak chip temperature by up to 17 °C compared to the state-of-the-art ML methods for an allowable performance overhead of only 5 .
more » « less
Full Text Available

Search for: All records